Papers by Kaushal Kumar Maurya

5 papers
LLMs cannot spot math errors, even when allowed to peek into the solution (2025.emnlp-main)

Copied to clipboard

Challenge: Large language models (LLMs) demonstrate impressive performance on existing reasoning benchmarks, but struggle with meta-reasoning tasks such as locating the first error step in student solutions.
Approach: They propose an approach that generates an intermediate corrected student solution, aligning more closely with the original student’s solution, which helps improve performance.
Outcome: The proposed approach generates an intermediate corrected student solution, aligning more closely with the original student’s solution, which helps improve performance.
SelectLLM: Query-Aware Efficient Selection Algorithm for Large Language Models (2025.findings-acl)

Copied to clipboard

Challenge: Existing large language models struggle with complex tasks such as factually-grounded reasoning and planning due to inherent training biases, model size constraints, and the quality or diversity of pre-training datasets.
Approach: They propose a novel algorithm to select the most suitable LLMs from a large pool and use it to efficiently generalize and perform tasks.
Outcome: The proposed model outperforms existing ensemble-based baselines and achieves competitive performance with similarly sized top-performing LLMs while maintaining efficiency.
AITutor-EvalKit: Exploring the Capabilities of AI Tutors (2026.eacl-demo)

Copied to clipboard

Challenge: Personalized one-on-one tutoring is an effective educational approach, yet its widespread adoption is constrained by the limited availability of qualified tutors and the high costs associated with tutor training.
Approach: They propose an evaluation tool that uses language technology to evaluate the pedagogical quality of AI tutors.
Outcome: The proposed evaluation tool is aimed at education stakeholders as well as the *ACL community at large, as it supports learning and can also collect user feedback and annotation.
ZmBART: An Unsupervised Cross-lingual Transfer Framework for Language Generation (2021.findings-acl)

Copied to clipboard

Challenge: Recent advances in NLP focus on large annotated training data.
Approach: They propose an unsupervised framework that does not use parallel or pseudo-parallel/back-translated data.
Outcome: The proposed framework does not use parallel or pseudo-parallel/back-translated data.
Unifying AI Tutor Evaluation: An Evaluation Taxonomy for Pedagogical Ability Assessment of LLM-Powered AI Tutors (2025.naacl-long)

Copied to clipboard

Challenge: Existing evaluations of large language models have been limited to subjective protocols and benchmarks.
Approach: They propose a unified evaluation taxonomy with eight pedagogical dimensions based on key learning sciences principles to assess the pedagical value of LLM-powered AI tutor responses grounded in student mistakes or confusions in the mathematical domain.
Outcome: The proposed taxonomy, benchmark, and human-annotated labels will streamline the evaluation process and help track the progress in AI tutors’ development.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations